Papers with VLMs process
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine (2025.naacl-long)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) have demonstrated promise in generating visually grounded responses, but their application in the medical domain is hindered by unique challenges. |
| Approach: | They propose a vision language model with versatile visual grounding for medicine that generates semantic segmentation masks and instance-level bounding boxes. |
| Outcome: | The proposed model can generate semantic segmentation masks and instance-level bounding boxes, and accommodates various imaging modalities, including both 2D and 3D data. |
Evaluating Vision-Language Models for Emotion Recognition (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Vision-Language Models (VLMs) have been used for objective multimodal reasoning tasks for decades. |
| Approach: | They present a comprehensive evaluation of large vision-language models for recognizing evoked emotions from images. |
| Outcome: | The proposed model performs well in evoked emotion recognition task and is robust to human errors. |
Finding Culture-Sensitive Neurons in Vision-Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Vision-language models struggle on culturally situated inputs, study shows . despite impressive performance, many VLMs struggle on such culturally grounded inputs . |
| Approach: | They propose a new margin-based selector to identify neurons associated with cultural selectivity . they also introduce a model-dependent decoder to identify such neurons . |
| Outcome: | The proposed model outperforms probability- and entropy-based methods in identifying neurons associated with cultural selectivity. |